Remote Data Jobs · Data Pipelines

Job listings

$151,000–$206,000/yr
US Canada Unlimited PTO

  • Build and improve machine learning models and data-driven systems that classify, cluster, label, and enrich Internet-observed assets and services.
  • Own the design and development of applied ML workflows that turn raw Internet telemetry into usable context for internal systems and customer-facing products.
  • Partner with engineering, research, security, and product teams to ensure we’re building the right models, datasets, and feedback loops.

Censys provides real-time Internet intelligence and actionable threat insights by mapping the Internet through IP scanning and web crawling. Founded by security researchers, it is a startup with midwestern roots that values diversity and inclusion.

$195,000–$270,000/yr

  • Lead end-to-end architecture of scalable Data Lakehouse solutions on GCP using BigQuery, GCS, and Dataplex.
  • Collaborate with customers to translate business goals into robust architectural blueprints and actionable plans.
  • Design and implement data pipelines for real-time and batch ingestion with modern orchestration frameworks.

Egen is a data-first technology company that helps clients drive action and impact through data and insights using advanced platforms like Google Cloud and Salesforce. It is a fast-growing, entrepreneurial company with a culture dedicated to learning, innovation, and solving tough problems.

Data Engineer

Ohr
$140,000–$200,000/yr

  • Build and maintain secure connectors across data platforms like Google Drive, Slack, and Claude.
  • Partner with scientists to integrate lab data collection tools into a unified, queryable platform.
  • Design pipelines for messy real-world data and own end-to-end infrastructure, from schema to monitoring.

Ohr creates molecules from atoms up using biocatalysis and primordial chemistry, powering rockets and securing industries with cleaner, scalable systems. As an early-stage company with a small, urgent team, it cultivates a culture of vision, hustle, and collaboration.

  • Design, develop, and optimize data pipelines to support reporting and analytics requirements.
  • Implement and manage reporting solutions using tools like Metabase, ensuring user-friendly dashboards and visualizations.
  • Collaborate with business stakeholders to gather requirements, define KPIs, and deliver actionable insights through reports.

SpryPoint is modernizing how utilities serve their communities with a cloud-native customer service and operations platform. Founded in 2011, they have grown to 300+ employees serving 100+ utility clients across North America and the Caribbean, with a mission to provide better technology to utilities.

  • Support training, fine-tuning, and evaluation of neural foundation models.
  • Build and maintain data pipelines for petascale neurobehavioural datasets.
  • Prototype tools and demos to apply trained models to robotics and embodied AI tasks.

Netholabs builds AI grounded in biological intelligence by recording petascale neurobehavioural data to train neural foundation models. They are a small, fast-moving research team focused on AI, robotics, and personalized intelligence.

  • Design and build scalable, fault-tolerant data pipelines for batch and streaming architectures.
  • Own data quality and observability end-to-end with monitoring, alerting, and automated testing.
  • Architect enterprise-grade datasets using appropriate modeling approaches like dimensional modeling or Data Vault.

The company builds foundational data infrastructure for analytics and AI at enterprise scale. The engineering team fosters a technically ambitious environment focused on long-term platform health and scalable architecture.

  • Lead end-to-end data engineering projects from discovery to production deployment, including complex data modeling and high-volume matching.
  • Design platform-agnostic data pipelines and schemas, ensuring scalability and performance across Snowflake, Databricks, and other platforms.
  • Define data integration strategies and communicate tradeoffs to clients, guiding architecture and tool selection.

Blend is a premier AI services provider that leverages data science, AI, and technology to drive client impact. The company fosters a culture of innovation and collaboration, working with global teams and clients.

  • Own the health of EULER's data end to end, from capture to validation to structure.
  • Build validation and reconciliation to ensure data integrity and trust.
  • Design canonical models and taxonomies for consistent data across customers.

EULER is an AI-native partner relationship management platform that helps companies capture and manage partnerships, which drive 30-50% of revenue. The company grew 600% last year and offers a remote, supportive culture where employees own their domains.